Papers with Penn Discourse Treebank

15 papers
Non-Topical Coherence in Social Talk: A Call for Dialogue Model Enrichment (2020.acl-srw)

Copied to clipboard

Challenge: Current models of dialogue focus on utterances within a topically coherent discourse segment, not on social conversations . a pilot annotation study of NTUs is a first step towards a model capable of rationalizing conversational coherence in social talk.
Approach: They conduct a pilot annotation study of social dialogues as a first step towards a Bayesian game-theoretic model . they first annotate content-based coherence relations that are not available in Disco-SPICE .
Outcome: The proposed model can rationalize conversational coherence in social talk, the authors say . the study focuses on the natural occurring social dialogues in the Disco-SPICE corpus .
CITE: A Corpus of Image-Text Discourse Relations (N19-1)

Copied to clipboard

Challenge: a crowd-sourced resource characterizes inferences in image-text contexts in the domain of cooking recipes . a recent study has found that image-image presentations are more effective at integrating text and image .
Approach: They propose a crowd-sourced resource for multimodal discourse characterizing inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations.
Outcome: The proposed corpus enables a better understanding of communication and common-sense reasoning . it is particularly important for automating the understanding and generation of text-image presentations .
The Role of Context and Uncertainty in Shallow Discourse Parsing (2022.coling-1)

Copied to clipboard

Challenge: Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning.
Approach: They hypothesize that context plays an important role in accurate human annotation and add uncertainty measures can improve model accuracy and calibration.
Outcome: The proposed model can be better calibrated by adding uncertainty measures to models with better accuracy and calibration.
Towards Identifying Alternative-Lexicalization Signals of Discourse Relations (2022.coling-1)

Copied to clipboard

Challenge: Existing shallow discourse parsing methods have been limited to identifying relations signaled by a discourse connective and those without a signal.
Approach: They propose to identify relations signalled by a discourse connective and those without . they compare a pattern-based approach and a sequence labeling model .
Outcome: The proposed approach is based on a pattern-based approach and a sequence labeling model.
Announcing the Prague Discourse Treebank 3.0 (2024.lrec-main)

Copied to clipboard

Challenge: PDiT 3.0 contains 21,662 discourse relations (plus 445 list relations) in 49 thousand sentences.
Approach: They present the Prague Discourse Treebank 3.0, a new version of the annotation of discourse relations marked by primary and secondary discourse connectives in the Prague Dependency Treebank.
Outcome: The new version of the PDiT 3.0 brings a largely revised annotation of discourse relations and achieves consistency with a Lexicon of Czech Discourse Connectives (CzeDLex) and sense taxonomy.
Entity Enhancement for Implicit Discourse Relation Classification in the Biomedical Domain (2021.acl-short)

Copied to clipboard

Challenge: Discourse relation classification is a challenging task when the text domain is different from the standard Penn Discourse Treebank (PDTB) training corpus domain.
Approach: They propose to use the Biomedical Discourse Relation Bank to improve discourse relational argument representation by linking explicit instances of similar relations with a voting pipeline.
Outcome: The proposed model outperforms the pre-trained BioBERT model by 2% points.
TED-CDB: A Large-Scale Chinese Discourse Relation Dataset on TED Talks (2020.emnlp-main)

Copied to clipboard

Challenge: TED-CDB dataset is a unique corpus of spoken discourse in Chinese . TED is based on the concept that discourse relations are grounded in an identifiable set of discourse connectives or Altlex expressions.
Approach: They have created a dataset that annotates TED talks in Chinese . they propose to adapt the dataset to Chinese news text to improve its performance .
Outcome: The TED-CDB dataset can improve the performance of systems for languages other than Chinese . it is adapted to features that are not present in English and can extract discourse semantic features .
The Causal News Corpus: Annotating Causal Relations in Event Sentences from News (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotation guidelines for event causality focus on only explicit relations or clauses.
Approach: They propose an annotation schema for event causality that addresses these concerns . they annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not.
Outcome: The proposed annotation schema for event causality addresses these concerns . it performs well with 81.20% F1 score on test set and 83.46% in 5-folds cross-validation .
Inducing Discourse Marker Inventories from Lexical Knowledge Graphs (2022.lrec-1)

Copied to clipboard

Challenge: Discourse marker inventories are important tools for the development of discourse parsers and corpora with discourse annotations.
Approach: They explore the potential of multilingual lexical knowledge graphs to induce multilingual discourse marker lexicons using concept propagation methods previously developed in translation inference across dictionaries.
Outcome: The proposed method can induce multilingual discourse marker lexicons using multilingual knowledge graphs.
Interactively-Propagative Attention Learning for Implicit Discourse Relation Recognition (2020.coling-main)

Copied to clipboard

Challenge: Existing models for discourse relation recognition use self-attention and interactive-attention mechanisms.
Approach: They develop a propagative attention learning model using a cross-coupled two-channel network.
Outcome: The proposed model improves on the baseline models on a Penn Discourse Treebank.
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)

Copied to clipboard

Challenge: a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task.
Approach: They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank.
Outcome: The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank.
Employing the Correspondence of Relations and Connectives to Identify Implicit Discourse Relations via Label Embeddings (P19-1)

Copied to clipboard

Challenge: Existing models for implicit discourse relation recognition lack the ability to accurately map connectives into discourse relations.
Approach: They propose a multi-task learning framework where relations and connectives are simultaneously predicted and leveraged to transfer knowledge between the two prediction tasks.
Outcome: The proposed framework yields state-of-the-art performance on several settings of the Penn Discourse Treebank dataset.
DisSent: Learning Sentence Representations from Explicit Discourse Relations (P19-1)

Copied to clipboard

Challenge: Existing models train on vast amounts of text or require costly, manually curated datasets.
Approach: They propose to leverage the discourse relations between sentences to curate a high quality sentence relation task by leveraging explicit discourse relations.
Outcome: The proposed model can be used to learn the meaning of two sentences in a bidirectional LSTM sentence encoder.
Multi-Label Classification for Implicit Discourse Relation Recognition (2024.findings-acl)

Copied to clipboard

Challenge: Prior research in discourse relation recognition has treated these instances as separate examples during training, with a gold-standard prediction matching one of the labels considered correct at test time.
Approach: They propose to use multiple labels to annotate an example when multiple relations are believed to hold simultaneously.
Outcome: The proposed frameworks don't depress performance for single-label prediction.
Discourse Sense Flows: Modelling the Rhetorical Style of Documents across Various Domains (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work on shallow discourse parsing has given renewed attention to the role of discourse relation signals, in particular explicit connectives and alternative lexicalizations.
Approach: They propose a model for extracting and classifying discourse relation signals from the Penn Discourse Treebank v3 corpus and introduce a new way of modeling rhetorical style by the linear order of coherence relations.
Outcome: The proposed models are based on the Penn Discourse Treebank v3 corpus and employ n-gram patterns to predict genre/domain discrimination.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations